Skip to main content

Custom API Language Model Requirements

What your application must provide to integrate as a Custom API Language Model with DynamoEval.

For background on why Custom Applications exist and how request/response adaptation works, see Custom Systems Overview. For formats, authentication UI, and SDK examples, see Custom API Language Models.

Required​

RequirementDetails
HTTP POST endpointA URL DynamoEval can call, including path (for example https://api.example.com/v1/chat).
REST JSON responseEach call must return a complete JSON body over REST. The following are not supported as the integration surface DynamoEval calls directly: SSE / text/event-stream, WebSockets, and gRPC. If your application only exposes one of those today, place a REST aggregation proxy in front that accepts JSON POST, talks to the upstream protocol, and returns a single JSON response—see Bridging Streaming APIs.
Fixed request and response shapeA stable schema for both request and response. Provide OpenAPI/Swagger, a sample curl, or example request and success-response JSON so JSONata transforms can be defined when your contract differs from DynamoEval's standard formats.
Long-lived authenticationOne of: No Authentication, Bearer Token, or API Key (including header name and optional scheme prefix when using API Key). Credentials must be usable on every evaluation request. Short-lived tokens are not supported—evaluations can run for hours, and an expired token causes the run to fail midway. The exception is an endpoint protected by Microsoft Entra ID, for which DynamoEval mints and refreshes tokens itself. See Authentication notes below.
Custom headersAny required headers beyond authentication (for example model-routing or tenant headers). State explicitly if none are required.
Network reachabilityThe endpoint must be reachable from the Dynamo deployment under test—over the public internet, or from within the same VPC / private network for private deployments.
Input size limitsAny maximum input length (characters or tokens) the application enforces.

Authentication notes​

DynamoEval attaches the configured credential to each evaluation call. Evaluations often run for hours and issue many requests over that period. If the token expires midway, later calls fail and the evaluation fails. That is why credentials must be a stable, long-lived secret (static Bearer token, API key, or no auth).

Not supported today

  • Short-lived access tokens (for example tokens that expire every few minutes), other than Microsoft Entra ID tokens that DynamoEval mints itself
  • Interactive login or browser-based OAuth flows
  • Auth that requires DynamoEval to refresh tokens on its own between requests, other than Microsoft Entra ID

If your application only issues short-lived tokens

Prefer one of these remediations before integration:

  1. Service account or machine credential — Issue a long-lived API key or service-user token dedicated to DynamoEval evaluation traffic (most common path).
  2. Auth-handling proxy — Run a small proxy that DynamoEval calls with a long-lived credential (or no auth). The proxy obtains or refreshes short-lived upstream tokens and forwards requests to your application. The same pattern used for streaming bridges applies—see Bridging Streaming APIs.

Microsoft Entra ID​

If your endpoint, or the Azure API Management gateway in front of it, accepts Microsoft Entra ID access tokens, DynamoEval can request and refresh those tokens itself through workload identity, so no proxy is needed. Your deployment must enable Microsoft Entra Workload Identity, and the endpoint must:

  • Accept tokens issued for the scope configured on the AI system.
  • Authorize the platform identity, or the customer-owned identity named on the AI system.
  • Be listed in the deployment's allowed endpoint hosts.

These are not required to register the model, but they help size evaluation concurrency and timeouts:

ItemDetails
ThroughputSustainable requests per second the application can handle.
Rate limitsPer-key or per-endpoint quotas, if any.
LatencyTypical and high-percentile (for example p95) response times.
Multi-turn supportWhether the API accepts conversation history (system / user / assistant turns).